Papers with mechanistic interventions
From Reasoning to Answer: Empirical, Attention-Based and Mechanistic Insights into Distilled DeepSeek R1 Models (2025.emnlp-main)
Copied to clipboard
| Challenge: | Large Reasoning Models generate explicit reasoning traces alongside final answers . the extent to which these traces influence answer generation remains unclear . |
| Approach: | They conduct empirical evaluation of Large Reasoning Models that include explicit reasoning . they also show that answer tokens attend substantially to reasoning tokens . |
| Outcome: | The results show that including explicit reasoning improves answer quality across domains . they also show that answer tokens attend substantially to reasoning tokens - the authors . |
How Do LLMs Generate Contrastive Sentiments? A Mechanistic Perspective (2026.eacl-long)
Copied to clipboard
| Challenge: | Despite extensive research, the mechanisms underlying LLMs' abilities remain poorly understood. |
| Approach: | They propose and validate a mechanistic intervention that transforms the sentiment of a text from positive to negative while making minimal edits. |
| Outcome: | The proposed intervention increases sentiment flip rate without sacrificing minimal changes to text content. |